Back

Translational Vision Science & Technology

Association for Research in Vision and Ophthalmology (ARVO)

Preprints posted in the last 7 days, ranked by how well they match Translational Vision Science & Technology's content profile, based on 39 papers previously published here. The average preprint has a 0.05% match score for this journal, so anything above that is already an above-average fit.

1
Optimizing Aqueous Humor Liquid Biopsy: Safety and Performance of a Short, Low-Dead-Space Ophthalmic Needle for Anterior Chamber Paracentesis

Singh, A. M.; Yeh, T.-C.; DeBoer, C.; Al-Moujahed, A.; Lin, J. B.; Smith, S. J.; Sanislo, S.; Janjua, K. A.; Lin, T.-C.; Almeida, D. R. P.; Mruthyunjaya, P.; Mahajan, V. B.

2026-09-02 ophthalmology 10.64898/2026.08.26.26361364 medRxiv
Top 0.1%
19.6%
Show abstract

Purpose: To evaluate the safety, procedural performance, sample recovery, and surgeon preference of an ophthalmic needle designed specifically for anterior chamber (AC) paracentesis. Methods: In this multicenter study, AC paracentesis was performed in clinic and operating-room settings using a 32-gauge x 4-mm needle with low dead space. The procedure was evaluated using a standardized physician survey. Prespecified outcomes included procedure-related adverse events (primary outcome), needle entry and handling, aspiration and sample recovery, comparative performance versus a 30-gauge needle, and physician preference for future use. Results: A total of 110 needle uses by eight surgeons were included. No ocular complications occurred, including lens or iris injury, hyphema, AC collapse, wound leak, hypotony, infection, or retinal complication, and no procedure required needle exchange or conversion to another device. Two technical events without ocular sequelae were noted, in which needle entry was partial thickness and did not reach the AC (1.8%; exact 95% CI, 0.2%-6.4%). Physicians rated needle entry, handling and sample recovery as good or excellent. Compared with a 30-gauge needle, the study needle was rated as at least comparable across all assessed domains. All surgeons rated it better or much better for intra-procedural safety and preferred it for future AC taps. Conclusions and Relevance: This short, 32-gauge low-dead-space ophthalmic needle demonstrated a favorable safety profile and was preferred over a 30-gauge needle by all surgeons. As aqueous humor liquid biopsy expands in clinical diagnostics and trials, an ophthalmic-specific needle design may help improve the consistency and safety of aqueous humor collection for molecular analysis and broader clinical use. Keywords: Anterior chamber paracentesis; Aqueous humor; Liquid biopsy; Low dead space; Ophthalmic needle

2
Spectral and melanopic dose calibration of consumer see-through extended-reality glasses for controlled retinal photostimulation

Gaidica, M.; Rosengart, M.

2026-08-31 ophthalmology 10.64898/2026.08.26.26361398 medRxiv
Top 0.1%
9.8%
Show abstract

Light reaching the retina is a primary regulator of human circadian physiology, acting largely through melanopsin-expressing retinal ganglion cells with peak short-wavelength sensitivity. Delivering known, repeatable retinal doses outside the laboratory is difficult because conventional light sources leave viewing geometry, gaze, and ambient conditions uncontrolled. Consumer extended-reality (XR) glasses fix a bright binocular display in constant geometry relative to the eye, but their suitability as calibrated photic stimulators has not been established. Here we validate a commercial micro-OLED XR display (VITURE Luma Ultra) for controlled retinal photostimulation. A purpose-built host application renders exact 8-bit RGB stimuli while independently controlling hardware brightness and logging all intensity-determining state; spectral radiance was measured at the retinal position of a 3D-printed phantom head with an open-source miniature spectroradiometer, anchored to absolute units by a luminance transfer calibration. The blue primary peaks at 461 nm (FWHM 43 nm), is spectrally invariant across a >10-fold intensity range, and at maximum output delivers an estimated 299 lx melanopic equivalent daylight illuminance, above consensus daytime recommendations, while remaining roughly two orders of magnitude below photobiological safety limits. The red primary is visually effective with minimal melanopic drive (melanopic DER 0.10), enabling spectrally shifted evening stimulation. Unlike the immersive virtual-reality headsets previously used for calibrated light delivery, the see-through form factor preserves the wearer's view of the surroundings--relevant for clinical monitoring in supervised settings such as the intensive care unit. These results show that consumer XR glasses can serve as a dose-calibrated platform for wearable photostimulation using an open-source measurement chain, and provide groundwork for application-layer dose-response studies.

3
GLP-1 Receptor Agonist Initiation and Anti-VEGF Treatment Frequency in Diabetic Macular Edema: an IRIS(R) Registry Cohort Study

Nagalamadaka, P.; Ross, C. J.; Gilbert, J. B.; Stillman, H.; Ghauri, S. Y.; Dutton, S. M.; Kearney, W.; Li, J. H.; Leong, A.; Singh, R. P.; Krzystolik, M. G.

2026-08-31 ophthalmology 10.64898/2026.08.29.26361426 medRxiv
Top 0.2%
5.0%
Show abstract

Purpose: To evaluate whether initiation of GLP-1 receptor agonists (GLP-1RAs) is associated with anti-VEGF treatment burden in type 2 diabetes patients with diabetic macular edema (DME) in the IRIS(R) Registry (Intelligent Research in Sight). Methods: Incident GLP-1RA initiators were matched 1:1 with controls via Mahalanobis distance matching (9,896 pairs; N=19,792) on sociodemographics, DME risk factors, and factors influencing GLP-1RA prescription including hypertension, obesity, chronic kidney disease. A longitudinal mixed-effects event-study model evaluated monthly anti-VEGF injection frequency over a 36-month window (12 months before through 24 months after initiation), adjusting for DME duration. Visual acuity (VA) and central subfield thickness (CST) were secondary outcomes. Results: Following GLP-1RA initiation, anti-VEGF injection trajectories did not significantly differ between the matched GLP-1RA and control cohorts (interaction coefficients -0.18 to 1.59, P>0.05). Likewise, no differences in VA were observed between cohorts (-0.05 to 0.04 logMAR, P>0.05) or CST (-14.12 to 33.58 {micro}m, P>0.05). Conclusion: In these matched cohorts, GLP-1RA initiation was not associated with the trajectory of anti-VEGF use or changes in VA or CST. Precis We used the American Academy of Ophthalmology IRIS(R) Registry (Intelligent Research in Sight) to identify patients with DME. In 19,792 matched patients, there was no significant reduction in injection frequency post GLP1-RA initiation and no significant change in VA or CST.

4
Glaucoma and Diabetes Mellitus: A Comparative Evaluation of Comorbid Effect on Tear Quantity among Patients in Owerri, Imo State, Nigeria.

Chukwuoha, C. M.; Ovenseri-Ogbomo, G.; Azuamah, Y. C.; Odimegwu, N. E.; Obioma-Elemba, J. E.; Ugwoke, G.; Nkeremuzor, E. C.; Eronini, Y.; Ikoro, N. C.; Esenwah, E. C.

2026-09-02 ophthalmology 10.64898/2026.08.30.26361782 medRxiv
Top 0.2%
3.3%
Show abstract

Abstract Objective: Glaucoma is a chronic disorder that impairs ocular health and may exacerbate ocular surface disease leading to tear film instability, dry eye symptoms and decreased quality of life. This study compared changes in tear quantity among glaucoma subjects living with and without diabetes mellitus, attending an eye clinic in Nigeria. Methods: A comparative cross sectional research design was used. 157 subjects which comprised 74 glaucoma subjects living with diabetes mellitus and 83 glaucoma subjects living without diabetes mellitus participated in the study. Tear quantity assessment included the Schirmer I test and tear meniscus height (TMH) measurement. Descriptive statistics, independent samples t-test and Chi-square test were used to examine the data at 0.05 level of significance. Results: Glaucoma subjects living with diabetes mellitus showed substantially decreased tear production (11.4 +/- 6.8 mm) compared with glaucoma subjects living without diabetes mellitus (19.6 +/- 9.6 mm; p < 0.001). Tear meniscus height in glaucoma subjects living with diabetes mellitus (0.8 +/- 0.3 mm) was significantly greater than in subjects living without diabetes mellitus (0.7 +/- 0.3 mm; p = 0.034). Conclusion: Diabetes mellitus dramatically deteriorates the ocular surface function in glaucoma subjects by decreasing tear production, altering the tear meniscus height and increasing the severity of ocular surface symptoms. Routine glaucoma care, especially in patients with diabetes mellitus, should include a full ocular surface evaluation including Schirmer I test, TBUT, TMH, and OSDI assessment to allow early detection and management of ocular surface disease, better treatment adherence, and improved visual outcomes. Keywords: Glaucoma, Diabetes Mellitus, Tear production, Tear Meniscus Height, Ocular Surface Disease.

5
What Matters Most: A Multi-Stakeholder Study of Outcome Domains in Lower-Limb Prosthesis Use

Ahmed, M. E.; Karlsson-Brown, S.; Koufaki, P.; Ahmadi, M.; Mico-Amigo, E. M.

2026-09-03 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361544 medRxiv
Top 0.3%
1.7%
Show abstract

Purpose: Lower-limb prosthesis use involves interacting physical, psychosocial, and device-related outcomes that may not be fully captured by conventional clinical assessment. This study aimed to develop and evaluate a stakeholder-informed framework of outcome domains relevant to meaningful everyday prosthesis use. Materials and Methods: A mixed-methods participatory design comprised a structured synthesis of selected clinically relevant content from five established patient-reported outcome measures; semi-structured interviews and importance and actionability ratings with 18 contributors (12 prosthesis users, four clinicians, and two industrial partners); and integration of the synthesis, qualitative, and rating findings. Interview records were analysed using reflexive thematic analysis, and ratings were analysed descriptively. Results: The resulting framework comprised four interrelated domains: Mobility, Physical Function, Psychosocial Wellbeing, and Prosthesis Experience. Mobility showed the clearest convergence across stakeholder perspectives. Prosthesis users showed the largest importance actionability gap for Prosthesis Experience (4.5 vs 3.0), whereas clinicians showed the largest gap for Psychosocial Wellbeing (5.0 vs 3.0). Interviews highlighted day-to-day variability in prosthesis use and the influence of confidence, fatigue, comfort, environmental conditions, social context, and device usability. Conclusions: Meaningful outcome assessment in prosthetic rehabilitation should extend beyond mobility alone to consider physical function, psychosocial wellbeing, and prosthesis experience within everyday contexts. The proposed framework provides a stakeholder-informed foundation for multidimensional outcome assessment in prosthetic rehabilitation.

6
Empowering adults to manage their hearing loss: assessing the benefits of user-controlled, smartphone-connected hearing aids.

Maidment, D. W.; Habib, A.; Gomez, R.; Benton, C.; Ferguson, M. A.

2026-09-03 otolaryngology 10.64898/2026.08.30.26361775 medRxiv
Top 0.4%
1.1%
Show abstract

The availability of hearing aids that can connect wirelessly to smartphone technologies via Bluetooth has grown exponentially in recent years. However, there is limited evidence assessing the benefits of user-adjustability afforded by these devices. This study aimed to assess the benefits of smartphone-connected hearing aids and an accompanying application (or app) in new and existing hearing aid users. In this single-centre, prospective, observational study, 44 adult hearing aid users (14 new and 30 existing) were recruited. Participants were fitted bilaterally with smartphone-connected hearing aids that could be adjusted by the user via an app. Self-reported outcome measures were collected at fitting and after seven-weeks of using the device in everyday life. For both new and existing hearing aid users, significant improvements in social participation, hearing-related fatigue, and hearing aid benefit and satisfaction were found. For existing hearing aid users, all outcomes were significantly better for the smartphone-connected hearing aids plus app in comparison to their existing hearing aids that did not connect to a smartphone, all with moderate-to-large clinical effect sizes (d> .6). User-controllability via the app was identified as the key benefit, and most participants (68%) reported that the app met their needs 'extremely' or 'very well'. These results suggest that, when used in conjunction with an app, smartphone-connected hearing aids can improve hearing outcomes due to greater user-controllability to improve listening. Thus, smartphone-connected hearing aids have the potential to facilitate patient-centred care, empowering the individual to successfully manage their hearing loss.

7
A single-session randomised crossover fNIRS study comparing three upper-limb mirror therapy task paradigms in healthy adults: a study protocol

Yang, T.; Wei, S.; Wang, Y.; Bai, D.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.28.26361691 medRxiv
Top 0.5%
0.9%
Show abstract

Background Mirror therapy (MT)-specifically paradigms using mirror visual feedback (MVF)-is widely used in neurorehabilitation; however, mechanistic implementations vary substantially in movement content, rhythmicity and attentional demands. This protocol describes an acute mechanistic, within-participant fNIRS screening study designed to compare three prespecified upper-limb mirror-therapy task paradigms and to quantify associated subjective experience after each condition in healthy adults during a single visit. Methods and analysis This is a single-centre, within-participant, randomised crossover study conducted at Wuhan Wuchang Hospital (Wuhan, China). Healthy adults aged 18-35 years will complete three task conditions once each in a counterbalanced order using a 3*3 Latin-square scheme: UMT1 (task-oriented rhythmic functional movement), UMT2 (open-ended free movement with auditory control), and UMT3 (non-functional rhythmic movement). fNIRS will be acquired using the NirSmart-6000A system during a standardised block design. The primary outcome is ROI-level HbO activation quantified as GLM-derived {beta} estimates within the prespecified primary ROIs (bilateral SM1/M1 and bilateral PMC). Secondary outcomes include ROI-level windowed {Delta}HbO (5-20 s post-onset relative to the immediately preceding rest; descriptive only), ROI-level {Delta}HbR, and post-condition subjective ratings (illusion, immersion, confusion and fatigue; 1-7 Likert). Condition effects will be analysed using linear mixed-effects models with fixed effects for condition and period and prespecified multiplicity-adjusted pairwise contrasts. Ethics and dissemination Ethics approval was obtained from the Ethics Committee of Wuchang Hospital Affiliated to Wuhan University of Science and Technology (Approval No.: 2025-112-01; approved on 2025-08-21). The study is expected to be minimal risk. Findings will be disseminated through publication of this protocol manuscript and subsequent results manuscripts and conference presentations. Trial registration number Chinese Clinical Trial Registry (ChiCTR2600116634). This study is conducted as a prespecified mechanistic sub-study under the overarching registered project.

8
Large language model-augmented implicit surgical video review

Zhang, Z.; Qadir, M. I.; Ramchand, R.; Belwadi, M.; Ball, R. P.; Konstantinopoulos, K.; Abbey, E. M.; Ernsberger, K. T.; Guzman, M. J.; Hendren, S.; Holcomb, B. K.; Robb, B. W.; Stankowski, T.; Waters, J. A.; Stefanidis, D.; Bilimoria, K. Y.; Mohanty, S.; Kolbinger, F. R.

2026-08-31 surgery 10.64898/2026.08.25.26361071 medRxiv
Top 0.6%
0.6%
Show abstract

Surgical video interpretation is a promising medical artificial intelligence application. However, no existing video annotation method preserves the spatiotemporal complexity of surgeon reasoning. Here we show that verbal reasoning and visual attention can be converted into structured, machine-actionable records of intraoperative behaviours. Our method decomposes transcribed verbal commentary into video-anchored semantic feedback chunks, which are classified via a large language model, with spatial grounding to surgical scenes via eyegaze or cursor tracking. We demonstrate method validity and scalability on structured and unstructured annotation tasks. For quality feedback on full-length colorectal procedures, the method reached near-human fidelity for chunking (mean cosine similarity: 0.95, SD: 0.01) and semantic classification across observations (mean Cohen's kappa: 0.71, SD: 0.07) and evaluative triggers (mean Cohen's kappa: 0.67, SD: 0.14), with excellent usability ratings. For structured critical view of safety assessment in laparoscopic cholecystectomy, implicit annotation yielded excellent agreement with explicit reviewer ratings (Cohen's kappa: 0.83, 0.49 and 0.81 across three criteria). We anticipate this method will advance surgical data science by enabling scalable construction of meaningfully annotated surgical video datasets.

9
Feasibility study of gait analysis using a new Wearable Force Plate

Sanz Morere, C. B.; Garrido-Lopez, G.; Hayase, M.; Rueda, J.; An, Q.; Shimoda, S.; Moreno, J. C.; Navarro, E.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.30.26361786 medRxiv
Top 0.7%
0.5%
Show abstract

Static force plates (FP) are the gold standard for measuring ground reaction forces (GRF) and computing joint moments through inverse dynamics in gait analysis. However, they are restricted to controlled environments, and the number of steps analyzed is limited by the plates embedded in the floor. To address these limitations, portable solutions such as sensorized insoles, socks, or shoes have emerged. Yet, creating wearable systems capable of measuring three-dimensional GRF in real-world conditions remains challenging. Current sensorized shoes often incorporate thick sensors (up to 2 cm), reducing usability and limiting their application in pathological populations or dynamic tasks like running. This study evaluates the usability of ShokacShoes, a novel sensorized shoe integrating three thin, three-dimensional force sensors, and explores its potential as a Wearable Force Plate (WFP). Eight healthy participants performed slow, natural, and fast walking using two insole configurations. Force and temporal metrics were derived from WFP and FP data. Results indicate that WFP enables accurate step segmentation and detects significant effects of speed and insole type on temporal and force metrics, confirming its reliability under different walking conditions. Comparisons with FP revealed differences in force metrics and signal morphology, though temporal parameters remained consistent. These results are likely due to sensor quantity and positioning. Thereby, ShokacShoes represent a valid solution capable of measuring three-dimensional forces within commercial footwear. Future work will focus on validating the applicability of a new version of ShokacShoes against gold-standard FP in a comprehensive validation study involving diverse real-world scenarios and pathological conditions.

10
Medial Plantar Nerve Shear Wave Elastography and Viscosity Imaging for Differentiating Mild from Moderate Diabetic Peripheral Neuropathy

Gao, X.; Li, Y.

2026-09-02 radiology and imaging 10.64898/2026.08.28.26361645 medRxiv
Top 1%
0.2%
Show abstract

Objective: To examine how medial plantar nerve shear wave speed (Cs) and viscosity coefficient (Vi) are associated with the severity of diabetic peripheral neuropathy (DPN), and to assess their ability to differentiate adjacent severity categories. Materials and Methods: Based on TCSS, the 113 patients with type 2 diabetes mellitus were assigned to the non-DPN (n = 33), mild DPN (n = 46), and moderate DPN (n = 34) groups. Medial plantar nerve Cs and Vi were measured using shear wave elastography and viscosity imaging. Receiver operating characteristic analysis evaluated Cs, Vi, and their logistic regression-based combination; areas under the curves (AUCs) were compared using DeLong tests. Results: Cs and Vi increased progressively across the three groups (both P < 0.001). For non-DPN versus mild DPN, the AUCs of Cs, Vi, and the combined model were 0.688 (95% CI, 0.604-0.772), 0.741 (0.660-0.822), and 0.745 (0.665-0.826), respectively, without significant pairwise differences. For mild versus moderate DPN, the corresponding AUCs were 0.707 (0.625-0.789), 0.794 (0.724-0.865), and 0.799 (0.731-0.867). The combined model outperformed Cs (P = 0.045), whereas Cs versus Vi and Vi versus the combined model did not differ significantly (P = 0.162 and 1.000, respectively). Conclusion: Medial plantar nerve Cs and Vi increased with DPN severity. Their combination improved discrimination between mild and moderate DPN compared with Cs alone but not with Vi alone. Quantitative medial plantar nerve viscoelastic assessment may complement clinical severity grading.

11
An interpretable, formally verified point-of-care ultrasound risk equation for difficult videolaryngoscopy: development and internal validation

Oyarzun-Silva, R. A.; Hernandez-Hernandez, P.; Fernandez-Vaquero, M. A.; De Luis-Cabezon, N.

2026-09-02 anesthesia 10.64898/2026.08.28.26361621 medRxiv
Top 1%
0.2%
Show abstract

Background. Videolaryngoscopy still requires adjuncts or hyperangulated rescue in a clinically important minority, and bedside screening discriminates modestly. Point-of-care ultrasound (POCUS) of the anterior airway is a promising alternative, but existing prediction models are opaque or assume a pre-specified functional form. We developed and internally validated a parsimonious, fully disclosed POCUS risk equation whose form is recovered from data and whose structural properties are machine-checked by formal proof - to our knowledge the first formally verified clinical risk predictor - following TRIPOD+AI 2024. Methods. In a prospective single-centre, single-operator cohort of 259 adults undergoing elective videolaryngoscopy (no-Easy airway 68/259, 26.3%), Sequentially Thresholded Least Squares with bootstrap stability selection (B=300) screened a 71-term library of nine POCUS features and retained a seven-term logistic equation; a two-term bootstrap-stable model was pre-specified as robustness analysis. Internal validation used 5x10 repeated cross-validation plus temporal and device hold-outs, with pre-specified overfitting and optimism assessments. Five behavioural properties of the deployed equation were machine-checked in Lean 4. Results. Two interactions met the |c|/sigma_c>2 stability criterion: skin-to-epiglottis x skin-to-hyoid-bone distance and tongue volume x sagittal tongue area. The seven-term equation reached a 5x10 cross-validated C-statistic of 0.966 (optimism-corrected 0.968) and held across temporal and device hold-outs (0.94-0.97). Calibration-in-the-large matched prevalence, with cross-validated slope 0.90 attenuating to 0.625 out-of-time; standard recalibration restored 0.92 without loss of discrimination. The pre-specified two-term robustness model reproduced this performance (C-statistic 0.964-0.968; events-per-parameter 34; shrinkage 0.99), confirming the result is not an artefact of the screening stage. Net benefit over a clinical baseline was positive across 10-50% thresholds. All five Lean 4 theorems compiled without sorry. Conclusions. A sparse, formally verified POCUS equation predicts difficult videolaryngoscopy with high internally validated discrimination and quantified, modest overfitting. Because the equation was developed in a single-operator cohort and its inputs are operator-dependent, external validation requires prior harmonisation of the measurement protocol and operator credentialing.

12
When medical credentials conflict with stated accuracy: A factorial study of source credibility and answer revision in medical LLM interactions

Wojcik, S.; Rulkiewicz, A.; Domienik-Karłowicz, J.

2026-09-01 health informatics 10.64898/2026.08.28.26361634 medRxiv
Top 1%
0.2%
Show abstract

Large language models perform well on medical examinations, but users routinely challenge their answers and invoke professional roles, and it is unclear what a system does when a medical credential and a stated task-specific accuracy point in opposite directions. In a factorial experiment on 480 items from four Polish specialty examination sets and three consumer large language model systems (ChatGPT, Claude, Gemini), each item and system received eleven independent conversations. Conditions crossed attributed source role (medical student, experienced specialist), stated prior accuracy on similar questions (2/10, 8/10) and suggestion correctness. The primary outcome was adoption of a prespecified incorrect option when the baseline answer matched the official key, comparing a specialist described as 2/10 with a student described as 8/10. Baseline agreement with the key was 87.2% across 15,683 analyzable conversations. The incorrect option was adopted more often from the specialist described as 2/10 than from the student described as 8/10 (10.2% vs. 7.6%; adjusted risk difference +2.82 percentage points, 95% CI +0.65 to +4.99). Estimates varied across the three systems and only one system-specific interval excluded zero. In a prespecified exploratory analysis with a shared eligibility rule, correct suggestions were adopted far more often than incorrect ones (risk difference +35.7 percentage points, 95% CI +30.8 to +40.7), indicating selective rather than indiscriminate compliance. An incorrect suggestion from a specialist with low stated accuracy was therefore slightly more influential than the same suggestion from a student with high stated accuracy, although the difference was modest and varied across systems. Agreement reached only after a user has disclosed a preferred answer should not automatically be treated as an independent second opinion, and medical large language model systems should be evaluated on how they revise answers after such disclosure, not solely on initial accuracy.

13
REINA: A Recognize-Then-Infer Wearable-to-App AI Framework for Breast Cancer Rehabilitation

Zhuang, Q.; Mou, C.; Liu, B.; Fu, M. R.; King, G. W.

2026-08-31 rehabilitation medicine and physical therapy 10.64898/2026.08.29.26361725 medRxiv
Top 1%
0.2%
Show abstract

Breast cancer survivors frequently experience upper-limb impairments, making continuous monitoring essential for effective rehabilitation. We propose REINA (Recognize-Then-Infer Wearable-to-App AI Framework), a two-stage deep-learning approach for remote monitoring of motor function during breast cancer rehabilitation using wearable-device data. Inertial measurement unit (IMU) signals from wearable devices are first used to recognize physical activities via supervised learning, followed by an activity-specific recurrent neural network (RNN) to infer corresponding electromyography (EMG) signals. REINA establishes reliable inference of neuromuscular activity from wearable IMU data, enabling real-time, cost-effective assessment of motor function recovery in real-world settings.

14
Spatial Geometry and Prevalence of Tunneling and Undermining in Pressure Ulcers

Frade, S.; Tunyiswa, Z.; Shin, M.; Dirks, R.

2026-09-01 dermatology 10.64898/2026.08.28.26361615 medRxiv
Top 1%
0.2%
Show abstract

Background: Pressure ulcers often develop complex three-dimensional morphologies that extend beyond the visible wound surface. Subsurface extensions such as tunneling and undermining create hidden cavities that complicate clinical assessment and wound management. Despite their clinical relevance, the prevalence and spatial characteristics of these subsurface wound morphologies have not been well characterized at scale. Methods: We performed a registry-based analysis using data from the LIFT-OFF Pressure Ulcer Registry, which captures longitudinal clinical documentation of pressure ulcers treated in routine care. The registry included approximately 18,000 patients with 32,000 documented pressure ulcers. Spatial characteristics of tunneling and undermining were analyzed using measurements recorded during routine wound assessments, including tract length, direction, and circumferential extent. Directional and circumferential distributions of subsurface defects were examined to characterize wound geometry. Results: Tunneling was present in 764 of 14,700 full-thickness pressure ulcers (5.2%), whereas undermining occurred in 2,293 wounds (15.6%). Tunneling tracts were typically short and exhibited directional clustering relative to the wound bed. In contrast, undermining demonstrated broader circumferential distributions and frequently involved larger subsurface separations beneath the wound margin. Both morphologies demonstrated distinct spatial patterns across anatomical locations and wound stages. Conclusion: Tunneling and undermining are common subsurface features of pressure ulcers and exhibit distinct spatial geometries. Whereas tunneling manifests as directional tract-like extensions, undermining more frequently produces circumferential tissue separation beneath wound margins. Improved characterization of subsurface wound architecture may enhance assessment of wound complexity and provide information not captured by surface measurements alone. Future studies should evaluate whether these features contribute to wound severity assessment, prognosis, and risk stratification.

15
AI Video Analysis of Psychomotor Performance in EMS Education: Agreement With Human Evaluators Across Three Skills

Otte, J. H.; Cartagena, A.

2026-08-31 medical education 10.64898/2026.08.26.26361437 medRxiv
Top 2%
0.1%
Show abstract

Background. A primary constraint on the capacity of EMS programs to meet industry demand is psychomotor instruction and verification, requiring direct observation of each student by a qualified evaluator. Whether AI video analysis can relieve it is untested; none has been applied to EMS skill examination or compared with human examiners. Objective. To quantify human EMS evaluator inter-rater reliability and evaluate an AI video-analysis platform against it. Methods. In a prospective, fully crossed study, five certified EMS evaluators and an AI platform independently scored identical video-recorded EMT performances of cervical collar application (n=15), bag-valve-mask (BVM) ventilation (n=14), and medical assessment (n=15) on dichotomous checklists with critical-failure criteria. Agreement was assessed at item, score, and decision levels using Fleiss' kappa, Krippendorff's alpha, Gwet's AC1, and ICC(2,1)/ICC(2,k). Results. Human item agreement was moderate (kappa 0.409 to 0.467), as was single-rater reliability (ICC(2,1) 0.539 to 0.694), against good panel reliability (ICC(2,k) 0.854 to 0.919). Recorded pass/fail agreement was fair (kappa 0.297 to 0.388) and critical-failure agreement near zero for two skills (kappa 0.028, 0.119). AI alignment tracked rubric observability rather than task complexity: r = 0.857 (collar, exceeding every human), -0.173 (BVM), 0.664 (medical), and it was most lenient on two skills. Conclusions. Human evaluators are an imperfect standard, especially on critical failures. The AI was a legitimate additional rater where checklist items were discrete and visually verifiable, but not where credit required judging continuous quantities such as ventilation rate, volume, or suction duration. Defensible uses are formative and archival, not summative. These results reflect an early, non-specialist configuration: a baseline, not a limit.

16
Mechanical elements related to the development of patellofemoral pain syndrome, pain intensity, and functional disability: A cross-sectional study

Yaghoubi, N.; Eghbali, M.; Soleimanifar, M.; Hashemirad, F.; Arab, A.

2026-08-31 rehabilitation medicine and physical therapy 10.64898/2026.08.26.26361453 medRxiv
Top 2%
0.1%
Show abstract

Background and purpose: Patellofemoral pain syndrome (PFPS) is a multifaceted condition where proximal, local, and distal factors may contribute to symptoms and limitations. How these factors collectively contribute to PFPS remains poorly understood. Therefore, this study compared proximal, local, and distal mechanical characteristics between individuals with and without PFPS and investigated their association with pain intensity and functional disability. Methods: Eighty participants were included: 40 individuals with unilateral or bilateral PFPS, 40 healthy controls. Isometric muscle strength of hip, trunk, and ankle was assessed using a handheld dynamometer. Joint alignment (Q-angle, rearfoot angle, pelvic tilt) and muscle flexibility (iliotibial band, hamstrings, quadriceps, gastrocnemius, and soleus) were measured using standard clinical techniques. Pain severity was assessed using a visual analog scale (VAS), and functional disability was evaluated using the Kujala score. Results: Individuals with PFPS showed reduced iliotibial band flexibility, decreased hamstring and soleus length, lower hip abductor strength, and greater anterior and lateral pelvic tilt (all p < 0.02). Multivariate analysis identified reduced iliotibial band flexibility (OR = 7.48) and greater anterior pelvic tilt (OR = 11.75) as independent associates of PFPS. Anterior pelvic tilt predicted pain severity, while anterior trunk muscle strength and Q-angle predicted disability. Discussion: Reduced iliotibial band flexibility and increased anterior pelvic tilt were independently associated with PFPS, while anterior pelvic tilt predicted pain severity and anterior trunk muscle strength and Q-angle predicted functional disability. Clinical assessment and rehabilitation of PFPS should therefore extend beyond the knee to include iliotibial band flexibility, pelvic alignment, and trunk muscle strength.

17
Evaluating GPT-4o Model Proficiency and Clinical Reasoning for Antimicrobial Stewardship in Dentistry

Dick, M.; Madathil, S.; Patel, A.; Kapoor, H. S.; Sharma, M.; D'Souza, Z.; Hameed, S.; Abu-Samak, M.; Najirad, A.; Dwairi, D.; Radaideh, O.; Nicolau, B.

2026-09-03 dentistry and oral medicine 10.64898/2026.09.01.26361980 medRxiv
Top 3%
0.1%
Show abstract

Objectives: Dentists prescribe approximately one in ten antibiotics worldwide, yet antimicrobial stewardship (AMS) remains underemphasized in dental education. Large language models (LLMs) may support AMS training, but their proficiency and clinical reasoning in this context remain unclear. We evaluated GPT-4o's accuracy and clinical reasoning on dental antibiotic prescribing questions, stratified by question difficulty. Methods: We assembled 125 multiple-choice questions on dental antibiotic prescribing from eight peer-reviewed studies (2017-2023). GPT-4o answered each question and generated a clinical justification. Accuracy was assessed against source-study answer keys and examined across difficulty quartiles. Justifications were evaluated using an adapted 12-axis human-evaluation framework assessing scientific consensus, extent and likelihood of harm, inappropriate and missing content, bias, and both correct and incorrect comprehension, retrieval, and reasoning. Prophylaxis-specific questions were analysed separately. Results: GPT-4o correctly answered 72% of questions. Accuracy remained relatively stable across difficulty quartiles (78%, 78%, 65%, 70%). Experts rated 95.4% of justifications positively across the 12 axes. Comprehension, retrieval, and reasoning each exceeded 96.2% positive ratings. Missing content was the main weakness (7.8%), and 7.1% of justifications showed a moderate-to-severe potential for harm. Performance on prophylaxis-specific questions (98.1%) exceeded non-prophylaxis questions (93.0%). Conclusions: GPT-4o demonstrated moderate-to-high proficiency and clinically defensible reasoning in dental antibiotic prescribing questions. However, residual risks indicate that it is not suitable for unsupervised clinical use but shows potential as a supervised AMS educational tool.

18
Sensorimotor effects of heatwrap and exercise in acute low back pain: results of a randomised controlled trial

Cote Picard, C.; Desgagnes, A.; Tittley, J.; Mailloux, C.; Perreault, K.; Mercier, C.; Dionne, C. E.; Roy, J.-S.; Masse-Alarie, H.

2026-09-02 rehabilitation medicine and physical therapy 10.64898/2026.08.31.26361841 medRxiv
Top 3%
0.1%
Show abstract

Background: Heatwrap is recommended for acute low back pain (ALBP), and previous research found heatwrap plus exercise more effective than each intervention alone. While recommended by clinical guidelines, their impact on mechanistic outcomes is unknown. This trial aimed to (i) assess immediate and short-term effects of heatwrap alone or combined with exercise, compared with a sham heatwrap, on pain sensitivity, lumbar muscle activity, current pain intensity, and trunk flexion range of motion, and (ii) explore whether changes in pain sensitivity and lumbar muscle activity are associated with changes in clinical symptoms from baseline to 1-week follow-up. Methods: A randomised controlled trial took place at a single research center. Of 315 individuals screened for eligibility, 99 adults with ALBP were recruited and assigned to one of three intervention groups: heatwrap plus exercise (n=34), heatwrap alone (n=33) or sham heatwrap (n=32). Interventions were applied for one hour at the first visit, and immediate effects were measured. Then, interventions were applied for 7 days, and short-term effects were measured at 1-week follow-up. Outcomes included pressure pain threshold, temporal summation of pain, flexion-relaxation ratio, trunk range of motion and current pain intensity. Results: Heatwrap and exercise did not produce greater effects over time than heatwrap alone or a sham heatwrap on all outcomes, and changes in sensorimotor outcomes at one week were not associated with changes in symptoms. Conclusions: Heatwrap and/or exercises did not influence specifically the potential sensorimotor mechanisms tested in individuals with ALBP. Trial registration: ClinicalTrials.gov; registration number: NCT03986047

19
A Multi-Agent Large Language Model Reasoning Engine for Early Detection of Pediatric Growth Disorders

Rabbani, N.; Mettner, J.; Lee, K.; Soto-Rivera, C. L.; Windberger, A.; Santiago, K.; Hatoun, J.; Correa, E. T.; Vernacchio, L.; Kohane, I.

2026-08-31 health informatics 10.64898/2026.08.28.26361655 medRxiv
Top 3%
0.0%
Show abstract

Routine childhood growth surveillance is a cornerstone of pediatric care. Growth pattern abnormalities are often early manifestations of chronic disease. Yet subtle abnormalities are frequently underrecognized, leading to diagnostic delays and avoidable morbidity. We introduce SPROUT (System for Pediatric Recognition Of Undiagnosed Trajectories), a generalized, multi-agent large language model (LLM) reasoning system designed to identify a broad spectrum of pediatric growth-related conditions from longitudinal electronic health records (EHRs) earlier than standard clinical practice. Using a large pediatric primary care EHR dataset, we developed and validated SPROUT as a two-stage system. First, a highly specific LLM screener flags concerning longitudinal growth patterns. Second, an Orchestrator module coordinates a multidisciplinary panel of LLM agents to generate a ranked differential diagnosis. To correct systemic reasoning errors, a Trainer module injects meta-knowledge into the panel via a dedicated "Learner" agent. Diagnostic capability was evaluated using a walk-forward, visit-by-visit simulation leading up to the diagnosis date. The SPROUT screener model achieved 98% (83/85) specificity and 28% (9/32) sensitivity on a gold-standard dataset of pediatric primary care patients when evaluated one year before the index date, and 100% specificity and 47% sensitivity when evaluated using longitudinal data up to the day of diagnosis. When applied to 300 control patients (i.e., healthy or undiagnosed), the screener flagged 15. Subsequent expert panel review confirmed high suspicion for undiagnosed pathology in 33% (5/15) of these cases. In chronological walk-forward validation on disease cases, the diagnostic engine identified conditions well before standard-of-care documentation. One year prior to clinical diagnosis, the system achieved sensitivities of 81% for type 1 diabetes mellitus, 56% for pituitary disorders, and 44% for celiac disease. The SPROUT multi-agent system demonstrates the ability to detect a significant portion of latent growth-related pediatric conditions months to years before current clinical standards while minimizing false positives. These results support its potential as a decision support tool for reducing diagnostic delays in pediatric care.

20
Multi-organ aging quantified from routine chest CT predicts chronic disease risk and mortality

Sato, J.; Salehjahromi, M.; Zafar, A.; Muneer, A.; Xu, X.; Zhu, E.; Vokes, N. I.; Cascone, T.; Le, X.; Altan, M.; Gardner, E. E.; Sheshadri, A.; Ostrin, E. J.; Salahudeen, A. A.; Li, T.; Merad, M.; Chaudhuri, A. A.; Gerber, D. E.; Kay, F. U.; Godoy, M. C. B.; Carter, B. W.; Shroff, G. S.; Byers, L. A.; Chung, C.; Jaffray, D.; Rice, D.; Liao, Z.; Chang, J. Y.; Vaporciyan, A. A.; Gibbons, D. L.; Wu, C. C.; Heymach, J. V.; Zhang, J.; Wu, J.

2026-08-31 radiology and imaging 10.64898/2026.08.26.26361434 medRxiv
Top 3%
0.0%
Show abstract

Biological aging occurs heterogeneously across individuals and organs. However, current measures of biological age incompletely capture organ-specific differences in health and disease risk. Because chest CT visualizes multiple thoracic organs, it offers an opportunity to quantify structural aging across organ systems. Here, we developed MOSAIC-Age, a framework characterizing eight organ-specific aging clocks on chest CT. The clocks were developed and validated using 9,971 CT scans from CT-RATE and MIDRC, and subsequently locked and applied to two independent prospective cohorts with 35,293 participants from the National Lung Screening Trial and Genetic Epidemiology of COPD study. CT-derived biological age gaps (BAGs) were examined in relation to lifestyle and socioeconomic factors, prevalent comorbidities, incident chronic diseases, and all-cause and cause-specific mortality. Higher BAGs, indicating organs that appeared older on CT than expected for their chronological age, were broadly associated with adverse health characteristics, chronic disease burden, and increased mortality risk. Multiple disease outcomes were associated with aging across several organs, whereas in multivariable analyses including all eight organ-specific BAGs, the remaining associations were more organ specific. A greater number of markedly older-appearing organs and a faster pace of aging were each associated with higher mortality. Together, these findings demonstrate that routine chest CT captures both shared and organ-specific patterns of biological aging and establish CT-derived organ aging as a quantitative imaging biomarker for assessing multi-organ health and long-term disease risk.